Papers with VLMs process

3 papers
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine (2025.naacl-long)

Copied to clipboard

Challenge: Vision Language Models (VLMs) have demonstrated promise in generating visually grounded responses, but their application in the medical domain is hindered by unique challenges.
Approach: They propose a vision language model with versatile visual grounding for medicine that generates semantic segmentation masks and instance-level bounding boxes.
Outcome: The proposed model can generate semantic segmentation masks and instance-level bounding boxes, and accommodates various imaging modalities, including both 2D and 3D data.
Evaluating Vision-Language Models for Emotion Recognition (2025.findings-naacl)

Copied to clipboard

Challenge: Large Vision-Language Models (VLMs) have been used for objective multimodal reasoning tasks for decades.
Approach: They present a comprehensive evaluation of large vision-language models for recognizing evoked emotions from images.
Outcome: The proposed model performs well in evoked emotion recognition task and is robust to human errors.
Finding Culture-Sensitive Neurons in Vision-Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Vision-language models struggle on culturally situated inputs, study shows . despite impressive performance, many VLMs struggle on such culturally grounded inputs .
Approach: They propose a new margin-based selector to identify neurons associated with cultural selectivity . they also introduce a model-dependent decoder to identify such neurons .
Outcome: The proposed model outperforms probability- and entropy-based methods in identifying neurons associated with cultural selectivity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations